Back

Gene Reports

Elsevier BV

All preprints, ranked by how well they match Gene Reports's content profile, based on 14 papers previously published here. The average preprint has a 0.02% match score for this journal, so anything above that is already an above-average fit. Older preprints may already have been published elsewhere.

1
Identification of Highly Frequent TP53 Germline Variations in a Cohort of Head and Neck Cancer Cases

Ajaz, S.; Muneer, R.; Siddiqa, A.; Memon, M. A.

2021-09-13 oncology 10.1101/2021.09.09.21263314 medRxiv
Top 0.1%
18.9%
Show abstract

TP53 is a tumour suppressor gene. Its inactivation plays a significant role in the molecular pathology of cancers. TP53 germline mutations increase the risk of developing multiple primary cancers. However, the role of alterations in TP53 germline DNA in head and neck cancers (HNCs) is not well-established. HNCs comprise one of the most frequent cancers in South Asia. The present discovery study reports the investigation of germline variations in the TP53 gene in a cohort of 30 HNC patients from Karachi, Pakistan. Blood samples were collected and genomic DNA was extracted from white blood cells. TP53 has 11 exons, where exon 1 is not transcribed. After quality control of DNA, amplification of seven selected exons along with their splice sites, two intronic regions (introns 2-3 and 3-4), and 3UTR were carried out. Sanger sequencing was carried out in order to identify germline variations. Comparison with wild type sequence revealed rs1642785 G>C (intron 2-3) variation in 63.2%, PIN3 duplication (rs17878362) in intron 3-4 in 94.7%, and rs1042522 G>C in exon 4 (p.R72P) in 66.6% of the cases. In 3UTR, 13.4% of the analyzed cases carried either one of two variants, i.e., 17:7669567_8delCA or 17:7669560C>G. The latter variations are reported for the first time in literature. In conclusion, we report three highly frequent germline variations and two newly discovered variations in 3UTR of TP53 germline DNA in HNC patients from Pakistan. These results shall contribute to delineating the genetic component of HNCs with potential translational implications.

2
An evolutionary analysis of the SARS-CoV-2 genomes from the countries in the same meridian

Mastriani, E.; Rakov, A. V.; Liu, S.-L.

2020-11-13 genomics 10.1101/2020.11.12.380816 medRxiv
Top 0.1%
18.7%
Show abstract

In the current study we analyzed the genomes of SARS-CoV-2 strains isolated from Italy, Sweden, Congo (countries in the same meridian) and Brazil, as outgroup country. Evolutionary analysis revealed codon 9628 under episodic selective pressure for all four countries, suggesting it as a key site for the virus evolution. Belonging to the P0DTD3 (Y14_SARS2) uncharacterized protein 14, further investigation has been conducted showing the codon mutation as responsible for the helical modification in the secondary structure. According to the predictions done, the codon is placed into the more ordered region of the gene (41-59) and close the area acting as transmembrane (54-67), suggesting its involvement into the attachment phase of the virus. The predicted structures of P0DTD3 mutated and not confirmed the importance of the codon to define the protein structure and the ontological analysis of the protein emphasized that the mutation enhances the binding probability.

3
In silico Study of the Uncertainly Significant VCP Variants reveal major Structural, Functional fluctuations leading to potential Disease-Based Associations

DAS, T.; ROYCHOWDHURY, S.; DAS, P.

2024-02-27 genomics 10.1101/2023.11.04.565542 medRxiv
Top 0.1%
15.2%
Show abstract

Valosin containing protein is involved in a plethora of crucial functions from proteostasis, stress granule clearance to genome maintenance and ubiquitination. Hence, mutations in VCP can lead to a plethora of fatal diseases like Amyotrophic Lateral Sclerosis, Inclusion body myopathy with Paget disease of bone and frontotemporal dementia type 1, Spastic paraplegia, Charcot-Marie-Tooth disease type 2Y, Dementia, and Osteitis Deformans to name a few. Studies on VCPs disease phenotype relationship, structural, and functional modifications of the proteins stability, conservation, molecular dynamics, and post-translational modifications havent been performed. This in silico study investigated the variants of VCP (R95C, R95G, A160P, R191P, R191Q) which have conflicting interpretations of pathogenicity which are often depreciated and lack data. Additionally, this study screens the study cohort and the fatal diseases linked to all these variants. Interestingly through various computational tools and disease-based population studies, it was found that these variants are often found in patients linked with fatal diseases. The protein-protein interaction showed UFD1 has a direct association with VCP. The physicochemical parameters showed that A160P had the highest fluctuations of all the variants. VCPs protein secondary structure, molecular dynamics simulations of RMSD, RMSF, RoG, hydrogen bonds, and solvent accessibility were all comparatively impacted due to the changes caused by the variants. Box-plot and Principal component analysis of the MD simulations visualized the changes in the wild type of the protein. Research in wet labs and screening of patient cohorts is necessary to further characterize the diseases linked with these variants. This can potentially lead to the identification of biomarkers for fatal rare diseases.

4
Wide variabilities identified among spike proteins of SARS Cov2 globally-dominant variant identified

Pal, A.; Pal, A.

2020-09-27 genomics 10.1101/2020.09.26.314385 medRxiv
Top 0.1%
12.9%
Show abstract

SARS Cov2 is a newly emerged virus causing pandemic with fatality and co-morbidity. The greatest limitations emerged is the lack of effective treatment and vaccination due to frequent mutations and reassortment of the virus, leading to evolvement of different strains. We identified a wide variability in the whole genome sequences as well as spike protein variants (responsible for binding with ACE2 receptor) of SARS Cov2 identified globally. Structural variations of spike proteins identified from representative countries from all the continents, seven of them have revealed genetically similar, may be regarded as the dominant type. Novel non-synonymous mutations as S247R, R408I, G612D, A930V and deletion detected at amino acid position 144. RMSD values ranging from 4.45 to 2.25 for the dominant variant spike1 with other spike proteins. This study is informative for future vaccine research and drug development with the dominant type.

5
Genetic Susceptibility Of Cytotoxic T Lymphocyte-Associated Antigen 4 Gene Polymorphism In The Onset Of Arthritis

Mukhtar, M.; Sheikh, N.; Suqaina, S.; Saleem, T.; Mehmood, R.; Khawar, M. B.

2021-04-29 genetic and genomic medicine 10.1101/2021.04.27.21255970 medRxiv
Top 0.1%
12.3%
Show abstract

Cytotoxic T lymphocyte-associated antigen 4 (CTLA-4) gene plays a vital role in the activation of T-cells as a down regulator. CTLA-4 gene polymorphisms have implicated a potential risk factor for autoimmune disorders like arthritis. Therefore the current study was designed to determine the association of CTLA-4 gene polymorphism in the onset of rheumatoid and osteoarthritis in Pakistani individuals. Genotyping was performed on 300 RA, 316 OA, and 412 control subjects by direct sequencing method as well as polymerase chain reaction-restriction fragment length polymorphism (PCR-RFLP) technique. It was observed that allelic and genotypic frequency of rs5742909, rs231775, rs4553808, rs733618, and rs3087243 were significantly varied among patients and controls and considered as a significant risk factor in the onset of RA as well as OA. However, no mutation was identified on the rs11571317 polymorphic site. Haplotype CAGTCA and CAG TCG act as a protectant against disease onset whereas CAACCG was significant in disease onset. Mutation on rs231775 polymorphic site lead to the change of threonine into alanine It was concluded that CTLA-4 gene polymorphism is a significant risk factor in the onset of RA as well as OA. Large scale survey is required for the screening of the genetic markers for pre-diagnosis of the disease. SUMMARY STATEMENTThe study summarized that CTLA-4 gene polymorphism plays a key role in the arthritis onset in Pakistani population.

6
Mutations in CSRP3/MLP, a Z-disc associated gene are functionally associated with dilated cardiomyopathy in Indian population

Giri, P.; Dixit, R.; Kumar, A.; Mohapatra, B.

2021-11-05 genetic and genomic medicine 10.1101/2021.11.05.21265852 medRxiv
Top 0.1%
12.1%
Show abstract

CSRP3 is a LIM domain containing protein, known to play an important role in cardiomyocyte development, differentiation and pathology. Mutations in CSRP3 gene are reported in both dilated and hypertrophic cardiomyopathy (DCM and HCM), however, the genotype-phenotype correlation still remains elusive. To investigate the pathogenic potential of CSRP3 variants in our DCM cohort, we have screened 100 DCM cases and 100 controls and identified 3 non-synonymous variations, of which two are missense variants viz., c.233 GGC>GTC, p.G78V; c.420 TGG>TGC, p.W140C, and the third one is a single nucleotide polymorphism (SNP) c.46 ACC>TCC, p.T16S. These variants were absent from 100 control individuals (200 chromosomes). In vitro functional analysis has revealed reduction of CSRP3 protein level in stably-transfected C2C12 cells with p.G78V or p.W140C variants. Immunostaining demonstrates both cytoplasmic and nuclear localization of the wild-type protein, however variants p.G78V and p.W140C cause obvious reduction in the cytoplasmic expression of CSRP3 protein which is more pronounced in case of p.W140C. Disarrayed actin cytoskeleton was also observed in mutants. Besides, the expression of target genes namely Ldb3, Myoz2, Tcap, Tnni3 and Ttn are also downregulated in response to these variants. GST-pulldown assay has also showed a diminished binding of CSRP3 protein with -Actinin due to both variants p.G78V and p.W140C. Both 2D, 3D-modeling have shown confirmational changes. Most in silico tools predict these variants as deleterious. Taken together, all these results suggest the impaired gene function due to these deleterious variants in CSRP3, implicating its possible disease causing role in DCM.

7
Type 2 Diabetes Mellitus pathogenesis and role of Peptidylglycine Alpha-Amidating Monooxygenase (PAM) Gene elaboration by In-silico Analysis

Mohsin, M.; Yusof, H. M.

2025-10-14 bioinformatics 10.1101/2025.10.12.681891 medRxiv
Top 0.1%
10.8%
Show abstract

Type 2 Diabetes Mellitus (T2DM) is caused by pancreatic beta cell failure and alpha cell dysfunction that are central to this disease pathophysiology. For T2DM, the Peptidylglycine Alpha-Amidating Monooxygenase (PAM) gene has susceptible locus as identified by Genome-wide association studies (GWAS) though underlying molecular mechanisms remained poorly characterised. This study have explored comprehensive in-silico characterization of the PAM gene potential role in T2DM. InterPro was employed to identify protein family and it was observed it belongs to Peptidylglycine alpha-hydroxylating monooxygenase/peptidyl-hydroxyglycine alpha-amidating lyase family (IR000720), consisting of two catalytic domains and nine active regions. Clinically significant 153 genetic variants including non-synonymous SNPs within the locus were identified after analysis of dbSNP through UCSC Genome Browser. Intrinsically disrupted large region (aa 290-495) was observed during protein disorder prediction by utilizing IUPred-3. Ensembl and NCBI BLAST were used to analyse evolutionary conservation sites. This analysis resulted in high sequence similarity with common model organisms Mus musculus and Rattus norvegicus with similarity 90% and 89% respectively. These computational findings suggest specific PAM variants likely to disrupt protein function, providing validated future direction for experimental studies to confirm PAM gene pivotal role in T2DM.

8
Association of EBV (Type 1 and 2) with Histopathological Outcomes in Breast Cancer in Pakistani Women

Ilyas, Y.; Khan, S.; Khan, N.

2021-11-19 cancer biology 10.1101/2021.11.16.468790 medRxiv
Top 0.1%
10.4%
Show abstract

IntroductionBreast cancer is one of the major and frequent tumors in the public health sector globally. The rising global prevalence of breast cancer has aroused attention in a viral etiology. Other than genetic and hormonal roles, viruses like Epstein - Barr virus (EBV) also participate in the development and advancement of breast cancer. AimThis study was conducted to detect the frequency of EBV genotypes in breast cancer patients and compare it with histopathological breast cancer changes. MethodsFormalin-fixed paraffin-embedded samples of breast cancer (N=60) ages ranged from 22-70 years were collected. EBV DNA was isolated, amplified, typed through PCR, and correlated with histopathological outcomes of breast cancer using SPSS software version 26. ResultsOur findings suggest that among breast cancer factors, Invasive ductal carcinoma (IDC) was the most common pathological pattern found among patients (90%), observed statistically significant (p= 0.01275). In regards to clinical staging, 8 (13.3 %) patients diagnosed with stage I, 39 (65 %) with stage II, and 13 (21.6 %) with stage III reported statistically significant association (p=0.0003). EBV DNA was detected in 68.3% (41/60) breast cancer patients, reported a statistically significant difference between the prevalence of EBV in breast cancer patients and normal samples (p = 0.001). Of 41 EBV-positive samples, 40 were EBV-1, while only 1 had EBV-2 infection (p < 0.001). No influence on cancer histology was observed. Regarding the association of breast cancer with EBV, histological type (P =0.209), tumor stage (P = 0.48), tumor grade (0.356), tumor sizes (p= 0.976), age (p= 0.1055), tumor laterality (p= 0.533) and ER/PR status (p=0.773) showed no significant association. ConclusionEBV-1 is prevalent in breast cancer patients and associated with IDC in the study area. For conclusive evidence, more studies are required based on a large sample size and by using more sensitive techniques.

9
Expansion of triplet nucleotide repeats in primates and other vertebrates: an evolutionary perspective

Murthy, S.; Mishra, R. K.

2023-04-06 genomics 10.1101/2023.04.05.535742 medRxiv
Top 0.1%
9.9%
Show abstract

Triplet nucleotide repeat (TNR) expansion has been linked to more than 40 inheritable neurological, neuromuscular and neurodegenerative disorders. Increase in copy number beyond a threshold causes further rapid expansion of the repeats, leading to instability and disease via gain/loss of function, toxic RNA products or chromosome instability. An analysis of these repeat regions across vertebrates shows that these repeats have consistently either arisen late or have increased in copy number in vertebrates, most significantly in primates and particularly in humans. Many of the known diseases have neurological basis, suggests positive selection of these repeats for neuronal function. Late occurrence of the diseases implicates a lack of negative selection. This evolutionary trade-off, a higher neuronal capability at the cost of disease susceptibility, is further supported by the observation that most of the genes associated with TNR expansion diseases have neuronal function.

10
Genome similarities between human-derived and mink-derived SARS-CoV-2 make mink a potential reservoir of the virus

Khalid, M.; Al-ebini, Y.

2022-05-30 genomics 10.1101/2022.05.29.493871 medRxiv
Top 0.1%
9.7%
Show abstract

The SARS-CoV-2 has RNA as the genome, which makes the virus more prone to mutations. Occasionally, mutations help a virus to cross the species barrier. The SARS-CoV-2 infection to humans and minks (Neovison vison) are examples of zoonotic spillover. Many studies have been published on the analysis of human-derived SARS-CoV-2, here we performed mutation analysis on the minks-derived SARS-CoV-2 genome sequences. We analyzed all available full-length mink derived SARS-CoV-2 genome sequences on GISAID (214 from Netherlands and 133 from Denmark). We found that the mutation pattern in the Netherlands and Denmark derived samples were different. Out of a total of 201 mutations, we found in this study, only 13 mutations were common in the Netherlands and Denmark derived samples. We found 4 mutations prevailed in the Netherlands and Denmark mink derived samples and these 4 mutations are also reported to prevail in human-derived SARS-CoV-2.

11
COVID-19 Variants Database: A repository for Human SARS-CoV-2 Polymorphism Data

Rakha, A.; Rasheed, H.; Batool, Z.; Akram, J.; Adnan, A.; Du, J.

2020-06-11 genomics 10.1101/2020.06.10.145292 medRxiv
Top 0.1%
9.7%
Show abstract

COVID-19 is a newly communicable disease with a catastrophe outbreak that affects all over the world. We retrieved about 8,781 nucleotide fragments and complete genomes of SARS-CoV-2 reported from sixty-four countries. The CoV-2 reference genome was obtained from the National Genomics Data Center (NGDC), GISAID, and NCBI Genbank. All the sequences were aligned against reference genomes using Clustal Omega and variants were called using in-house built Python script. We intend to establish a user-friendly online resource to visualize the variants in the viral genome along with the Primer Infopedia. After analyzing and filtering the data globally, it was made available to the public. The detail of data available to the public includes mutations from 5688 SARS-CoV-2 sequences curated from 91 regions. This database incorporated 39920 mutations over 3990 unique positions. According to the translational impact, these mutations include 11829 synonymous mutations including 681 synonymous frameshifts and 21701 nonsynonymous mutations including 10 nonsynonymous frameshifts. Development of SARS-CoV-2 mutation genome browsers is a fundamental step obliging towards the virus surveillance, viral detection, and development of vaccine and therapeutic drugs. The SARS-COV-2 mutation browser is available at http://covid-19.dnageography.com.

12
Trends in Arabidopsis Research post genome sequencing- A Scientometric study

Kumar, S.; Kushwaha, A. K.; R, T.; M, B.; P, K.

2022-04-06 plant biology 10.1101/2022.02.22.481563 medRxiv
Top 0.1%
9.7%
Show abstract

1.Arabidopsis thaliana, a model plant, is intensively researched because of the intrinsic advantages associated with its life cycle, genetics, and other characteristics. In the present study, we analysed the publication data of research work done on A. thaliana during 2001-2020, a period when the whole genome sequence of the plant was available to the researchers. The current meta-analysis showed that 31965 research papers were published globally on Arabidopsis during the study period. The USA topped the countries with the maximum share of publications (28.97%), followed by China, Japan, and West European countries. The analysis showed a temporal shift in the number of publications from different countries and their research focus. After 2013, the research output from China was higher than that from the USA, and there was a shift in focus from developmental biology to stress-related topics. Plant Physiology journal carried the most research work (2490 publications) on Arabidopsis, which was followed by Plant Cell (2376), the Plant Journal (2260) and the Journal of Experimental Botany (1121). However, there was a progressive decline in the publication in these journals, and a shift was evident in favour of open access journals like Frontiers in Plant Science. The publication and citation numbers also show ongoing Arabidopsis researchs relevance to plant sciences, particularly curiosity-driven and discovery-based science. This article delves into the patterns in the prominence of research areas, ideas and foci for the future Arabidopsis research roadmap.

13
In silico analysis of single nucleotide polymorphisms (SNPs) in human C-C chemokine receptor type five (CCR5) gene

Ali Hassan, A.; Ibrahim, M. E.

2020-11-16 bioinformatics 10.1101/2020.11.14.382739 medRxiv
Top 0.1%
9.6%
Show abstract

IntroductionChemokines are small transmembrane proteins with immune surveillance and immune cell recruitment functions. the expression of CCR5 gene affects virus production and viral load(1). The CCR5 gene contains two introns, three exons, and two promoters, and it is necessary as a co-receptor for the entry of the macrophage-tropic HIV strains. Mutations in the coding region of CCR5 affect the protein structure, which will affect production, chemokine binding, transport, signaling and expression of the CCR5 receptor. MethodsSNPs within CCR5 gene were retrieved from ensemble database. Coding SNPs were analyzed using SNPnexus. Coding non-synonymous SNPs in CCR5 binding domains with Viral gp120 were analyzed using SIFT, PolyPhen and I-mutant tools. Project HOPE then used to modelled the 3D structure of the protein resulting from these SNPs. Non-coding SNPs that affects miRNAs in 3 rejoin were analyzing using PolymiRTS. SNPs that affect transcription factor binding were analyzed using regulomeDB. Results(178) non-synonyms missense SNPs were found to have deleterious and damaging effect on the structure and function of the protein. In CCR5 binding domains with Viral gp120: 3 SNPs rs145061115, rs199824195 and rs201797884 were found to affect both structure and function and stability of chemokine protein. The 2 SNPs rs185691679 and rs199722070 has a role in disruption and creation of the target sites in miRNA seeds due to their high conservation score. ConclusionMutations in CCR5 gene may explain and represent the molecular basis of the resistance to HIV infection.

14
In silico Proteome analysis of Severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2)

Baruah, C.; Devi, P.; Sharma, D. K.

2020-05-28 bioinformatics 10.1101/2020.05.23.104919 medRxiv
Top 0.1%
9.6%
Show abstract

Severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) (2019-nCoV), is a positive-sense, single-stranded RNA coronavirus. The virus is the causative agent of coronavirus disease 2019 (COVID-19) and is contagious through human-to-human transmission. The present study reports sequence analysis, complete coordinate tertiary structure prediction and in silico sequence-based and structure-based functional characterization of full SARS-CoV-2 proteome based on the NCBI reference sequence NC_045512 (29903 bp ss-RNA) which is identical to GenBank entry MN908947 and MT415321. The proteome includes 12 major proteins namely orf1ab polyprotein (includes 15 proteins), surface glycoprotein, ORF3a protein, envelope protein, membrane glycoprotein, ORF6 protein, ORF7a protein, orf7b, ORF8, Nucleocapsid phosphoprotein and ORF10 protein. Each protein of orf1ab polyprotein group has been studied separately. A total of 25 polypeptides have been analyzed out of which 15 proteins are not yet having experimental structures and only 10 are having experimental structures with known PDB IDs. Out of 15 newly predicted structures six (6) were predicted using comparative modeling and nine (09) proteins having no significant similarity with so far available PDB structures were modeled using ab-initio modeling. Structure verification using recent tools QMEANDisCo 4.0.0 and ProQ3 for global and local (per-residue) quality estimates indicate that the all-atom model of tertiary structure of high quality and may be useful for structure-based drug designing targets. The study has identified nine major targets (spike protein, envelop protein, membrane protein, nucleocapsid protein, 2-O-ribose methyltransferase, endoRNAse, 3-to-5 exonuclease, RNA-dependent RNA polymerase and helicase) for which drug design targets could be considered. There are other 16 nonstructural proteins (NSPs), which may also be percieved from the drug design angle. The protein structures have been deposited to ModelArchive. Tunnel analysis revealed the presence of large number of tunnels in NSP3, ORF 6 protein and membrane glycoprotein indicating a large number of transport pathways for small ligands influencing their reactivity.

15
In-silico Analysis of SARS-Cov2 Spike Proteins of Different Field Variants

Tariq, M. H.; Amir, A.; Ikram, A.

2023-01-23 bioinformatics 10.1101/2023.01.22.525048 medRxiv
Top 0.1%
9.6%
Show abstract

BackgroundCoronaviruses belong to the group of RNA family of viruses which trigger diseases in birds, humans, and mammals, which can cause respiratory tract infections. The COVID-19 pandemic has badly affected every part of the world, and the situation in the world is getting worse with the emergence of novel variants. Our study aims to explore the genome of SARS-,CoV2 followed by in silico analysis of its proteins. MethodsDifferent nucleotide and protein variants of SARS-Cov2 were retrieved from NCBI. Contigs & consensus sequences were developed to identify variations in these variants by using SnapGene. Data of variants that significantly differ from each other was run through Predict Protein software to understand changes produced in protein structure The SOPMA web server was used to predict the secondary structure of proteins. Tertiary structure details of selected proteins were analyzed using the online web server SWISS-MODEL. FindingsSequencing results shows numerous single nucleotide polymorphisms in surface glycoprotein, nucleocapsid, ORF1a, and ORF1ab polyprotein. While envelope, membrane, ORF3a, ORF6, ORF7a, ORF8, and ORF10 genes have no or few SNPs. Contigs were mto identifyn of variations in Alpha & Delta Variant of SARs-CoV-2 with reference strain (Wuhan). The secondary structures of SARs-CoV-2 proteins were predicted by using sopma software & were further compared with reference strain of SARS-CoV-2 (Wuhan) proteins. The tertiary structure details of only spike proteins were analyzed through the SWISS-MODEL and Ramachandran plot. By Swiss-model, a comparison of the tertiary structure model of SARS-COV-2 spike protein of Alpha & Delta Variant was made with reference strain (Wuhan). Alpha & Delta Variant of SARs-CoV-2 isolates submitted in GISAID from Pakistan with changes in structural and nonstructural proteins were compared with reference strain & 3D structure mapping of spike glycoprotein and mutations in amino acid were seen. ConclusionThe surprising increased rate of SARS-CoV-2 transmission has forced numerous countries to impose a total lockdown due to an unusual occurrence. In this research, we employed in silico computational tools to analyze SARS-CoV-2 genomes worldwide to detect vital variations in structural proteins and dynamic changes in all SARS-CoV-2 proteins, mainly spike proteins, produced due to many mutations. Our analysis revealed substantial differences in functional, immunological, physicochemical, & structural variations in SARS-CoV-2 isolates. However real impact of these SNPs can only be determined further by experiments. Our results can aid in vivo and in vitro experiments in the future.

16
Phylogenomic analysis of SARS-CoV-2 genomes from western India reveals unique linked mutations

Paul, D.; Jani, K.; Kumar, J.; Chauhan, R.; Seshadri, V.; Lal, G.; Karyakarte, R.; Joshi, S.; Tambe, M.; Sen, S.; Karade, S.; Anand, K. B.; Shergill, S. P. S.; Gupta, R. M.; Bhat, M. K.; Sahu, A.; Shouche, Y. S.

2020-07-31 genomics 10.1101/2020.07.30.228460 medRxiv
Top 0.1%
9.5%
Show abstract

India has become the third worst-hit nation by the COVID-19 pandemic caused by the SARS-CoV-2 virus. Here, we investigated the molecular, phylogenomic, and evolutionary dynamics of SARS-CoV-2 in western India, the most affected region of the country. A total of 90 genomes were sequenced. Four nucleotide variants, namely C241T, C3037T, C14408T (Pro4715Leu), and A23403G (Asp614Gly), located at 5UTR, Orf1a, Orf1b, and Spike protein regions of the genome, respectively, were predominant and ubiquitous (90%). Phylogenetic analysis of the genomes revealed four distinct clusters, formed owing to different variants. The major cluster (cluster 4) is distinguished by mutations C313T, C5700A, G28881A are unique patterns and observed in 45% of samples. We thus report a newly emerging pattern of linked mutations. The predominance of these linked mutations suggests that they are likely a part of the viral fitness landscape. A novel and distinct pattern of mutations in the viral strains of each of the districts was observed. The Satara district viral strains showed mutations primarily at the 3' end of the genome, while Nashik district viral strains displayed mutations at the 5' end of the genome. Characterization of Pune strains showed that a novel variant has overtaken the other strains. Examination of the frequency of three mutations i.e., C313T, C5700A, G28881A in symptomatic versus asymptomatic patients indicated an increased occurrence in symptomatic cases, which is more prominent in females. The age-wise specific pattern of mutation is observed. Mutations C18877T, G20326A, G24794T, G25563T, G26152T, and C26735T are found in more than 30% study samples in the age group of 10-25. Intriguingly, these mutations are not detected in the higher age range 61-80. These findings portray the prevalence of unique linked mutations in SARS-CoV-2 in western India and their prevalence in symptomatic patients. ImportanceElucidation of the SARS-CoV-2 mutational landscape within a specific geographical location, and its relationship with age and symptoms, is essential to understand its local transmission dynamics and control. Here we present the first comprehensive study on genome and mutation pattern analysis of SARS-CoV-2 from the western part of India, the worst affected region by the pandemic. Our analysis revealed three unique linked mutations, which are prevalent in most of the sequences studied. These may serve as a molecular marker to track the spread of this viral variant to different places.

17
Association between methylenetetrahydrofolate reductase gene C677T polymorphism and susceptibility to polycystic ovary syndrome

Rai, V.; Kumar, P.

2020-06-19 genetic and genomic medicine 10.1101/2020.06.15.20132324 medRxiv
Top 0.1%
9.3%
Show abstract

Polycystic ovary syndrome (PCOS) is the most common form of endocrinopathy of women. Several studies have investigated the association of methylenetetrahydrofolate reductase (MTHFR) gene C677T polymorphism with PCOS risk but the results are contradictory. So, the aim of the present study was to carry out a meta-analysis of a published case control studies to find out exact association between MTHFR gene C677T polymorphism and PCOS susceptibility. Pubmed, Springer link, Science Direct and Google Scholar databases were searched for case-control studies. Odds ratios (ORs) with 95% confidence intervals (CIs) was used as association measure and meta-analysis was performed using MIX and MetaAnalyst programs. Meta-analysis of 24 studies showed strong significant association between C677T polymorphism and PCOS risk (for T vs. C: OR= 1.18, 95% CI=1.01-1.38, p=0.03; for TT vs. CC: OR= 1.37, 95% CI=1.0-1.89, p= 0.045; for TT + CT vs. CC: OR= 1.31, 95% CI= 1.07-1.62, p= 0.008; for CT vs. CC: OR= 1.31, 95% CI= 1.04-1.62, p= 0.01 and for TT vs. CT + CC: OR= 1.10, 95% CI= 0.82-1.47, p= 0.04). In subgroup analysis, MTHFR C677T polymorphism is significantly associated with PCOS risk with Asian individuallas but in Caucasian population MTHFR C677T polymorphism was not significantly associated with PCOS risk. In conclusion, C677T polymorphism is a risk factor for PCOS.

18
The clinical feature of triple-negative breast cancer in Beijing, China

Zhao, H.; Feng, Y.; Yang, J.

2021-08-05 oncology 10.1101/2021.08.03.21261573 medRxiv
Top 0.1%
8.9%
Show abstract

ObjectiveTo collect and analyze the clinical feature of triple-negative breast cancer (TNBC) in Beijing, to provide the information for the local oncologist to make sound treatment plans. MethodThe clinical data of 280 breast cancer patients with TNBC admitted to the oncology hospital of China Academy of Medical Sciences were collected and divided into (recurrence and metastasis) group and non-(recurrence and metastasis) group. Breast cancer patients with TNBC were classified according to age distribution, family history of breast cancer, pathological type, histological grade, clinical stage, tumor thrombus, tumor size and lymph node metastasis. and 15 BRCA1 gene SNP loci were also measured by a high throughput Mass ARRAY time-of-flight mass spectrometry biochip system and compared the expression of 15 BRCA1 gene SNP loci between (recurrence and metastasis) group and non-(recurrence and metastasis) group. {chi}2 test was used to analyze the difference between two groups, and P<0.05 considered statistically significant. ResultsA total of 280 breast cancer patients with TNBC were enrolled in this study, median age 45 years old. There were 117 cases breast cancer patients with TNBC in (recurrence and metastasis) group, accounting for 41.79% in total breast cancer patients with TNBC and 163 cases breast cancer patients with TNBC in non-(recurrence and metastasis) group, accounting for 58.21% in total breast cancer patients with TNBC; There was no significant difference in age distribution, family history of breast cancer, pathological type and histological grade between non-(recurrence and metastasis) group and (recurrence and metastasis) group (P>0.05); but there were significant differences in clinical stage, vascular tumor thrombus, tumor size and lymph node metastasis between two groups (P<0.05); and then we compared the expression of 15 BRCA1 gene SNP loci in (recurrence and metastasis) group and non-(recurrence and metastasis) group, found that BRCA1gene rs 12516 CC loci (38.8% VS 44.4%), BRCA1gene rs 12516 TT loci (15.6% VS 10.4%), BRCA1 gene rs 16940 CC loci (15.1% VS 10.4%), BRCA1 gene rs 16940 TT loci (39.0% VS 44.8%), BRCA1 gene rs 16941 AA loci (38.1% VS 44%), BRCA1 gene rs 16941 GG loci (15.0% VS 10.3%), BRCA1 gene rs16942 AA loci (39.0% VS 44.8%), BRCA1 gene rs16942 GG loci (15.1% VS 10.4%), BRCA1gene rs799906 CC loci (15.9% VS 10.4%), BRCA1gene rs799906 TT loci (38.7% VS 44.8%), BRCA1gene rs799917 CC loci (38.7% VS 44.4%), BRCA1gene rs799917 TT loci (15.7% VS 10.4%), BRCA1gene rs1060915 CC loci (15.5% VS 10.4%), BRCA1gene rs1060915 TT loci (39.1%VS 44.8%), BRCA1gene rs1799966 AA loci (37.7% VS 44.4%), BRCA1gene rs1799966 GG loci (15.1% VS 10.4%), BRCA1 Gene rs2070833 AA loci (3.1% VS 7.0%), BRCA1 Gene rs2070833 CC loci (56.3% VS 51.3%), BRCA1gene rs3737559 GG loci(78.5% VS 84.5%), BRCA1gene rs3737559 GA loci(19.0% VS 14.6%), BRCA1gene rs8176199 AA loci (60.5% VS 64.6%), BRCA1gene rs8176318 GG loci (38.4% VS 43.4%), BRCA1gene rs8176318 TT loci (15.1% VS 10.6%), BRCA1gene rs8176323 CC loci (38.6% VS 43.9%), BRCA1gene rs8176323 GG loci (15.2% VS10.5%), BRCA1gene rs11655505 AA loci (14.9% VS 10.4%), BRCA1gene rs11655505 GG loci (39.1% VS 44.8%) had a difference at the accident rate between recurrence and metastasis group and non-(recurrence and metastasis) group, but the frequencies of genotypes in the (recurrence and metastasis) group and non-(recurrence and metastasis) group were similar, there was no statistical significant correlation between the SNP genotype of the BRCA1 gene and the recurrence and metastasis risk of TNBC (P>0.05). ConclusionsThere were higher recurrence and metastasis (41.79%) in total 280 cases breast cancer patients with TNBC in Beijing area; breast cancer patients with TNBC in Beijing area had a unique clinical feature no matter at clinical stage, vascular tumor thrombus, tumor size and lymph node metastasis or the expression of 15 BRCA1 gene SNP loci, those data may provide some information for clinical staff for TBNC treatment.

19
Severe Acute Respiratory Syndrome Coronavirus-2 genome sequence variations relate to morbidity and mortality in Coronavirus Disease-19

Mehta, P.; Sarkar, S.; Ghoshal, U.; Pandey, A.; Singh, R.; Singh, D.; Vishvkarma, R.; Ghoshal, U. C.; Maurya, R.; Pandey, R.; Ramachandran, R.; Bhadury, P.; Kundu, T. K.; Singh, R.

2021-05-24 genomics 10.1101/2021.05.24.445374 medRxiv
Top 0.1%
8.9%
Show abstract

Outcome of infection with Severe Acute Respiratory Syndrome Coronavirus-2 (SARS-CoV-2) may depend on the host, virus or the host-virus interaction-related factors. Complete SARS-CoV-2 genome was sequenced using Illumina and Nanopore platforms from naso-/oro-pharyngeal ribonucleic acid (RNA) specimens from COVID-19 patients of varying severity and outcomes, including patients with mild upper respiratory symptoms (n=35), severe disease ad-mitted to intensive care with respiratory and gastrointestinal symptoms (n=21), fatal COVID-19 outcome (n=17) and asymptomatic (n=42). Of a number of genome variants observed, p.16L>L (Nsp1), p.39C>C (Nsp3), p.57Q>H (ORF3a), p.71Y>Y (Membrane glycoprotein), p.194S>L (Nucleocapsid protein) were observed in similar frequencies in different patient subgroups. However, seventeen other variants were observed only in symptomatic patients with severe and fatal COVID-19. Out of the latter, one was in the 5UTR (g.241C>T), eight were synonymous (p.14V>V and p.92L>L in Nsp1 protein, p.226D>D, p.253V>V, and p.305N>N in Nsp3, p.34G>G and p.79C>C in Nsp10 protein, p.789Y>Y in Spike protein), and eight were non-synonymous (p.106P>S, p.157V>F and p.159A>V in Nsp2, p.1197S>R and p.1198T>K in Nsp3, p.97A>V in RdRp, p.614D>G in Spike protein, p.13P>L in nucleocapsid). These were completely absent in the asymptomatic group. SARS-CoV-2 genome variations have a significant impact on COVID-19 presentation, severity and outcome.

20
First Report of Evaluation of Variant rs11190870 nearby LBX1 Gene with Adolescent Idiopathic Scoliosis Susceptibility in a South-Asian Indian Population.

Singh, H.; Shipra, S.; Gupta, M.; Bhau, P.; Chalotra, T.; Manotra, R.; Gupta, N.; Gupta, G.; Pandita, A. K.; Butt, M. F.; Sharma, R.; Pandita, S.; Singh, V.; Garg, B.; Rai, E.; Sharma, S.

2022-06-28 genetic and genomic medicine 10.1101/2022.06.28.22276987 medRxiv
Top 0.1%
8.3%
Show abstract

LBX1 is a developmental gene involved in skeletal muscle development and somatosensory functioning and proven to be an important gene involved in Adolescent Idiopathic Scoliosis (AIS) etiology. Variant rs11190870 is located 7.5 kb downstream of LBX1 gene and is part of haplotype that is reported to provide risk for AIS. Several studies, including various Genome Wide Association, replication and meta-analyses studies have implicated its association with AIS in different populations. However, any such study is altogether lacking in South-Asian Indian populations. In this first genetic association study for AIS from the region, we tried to replicate association of variant rs11190870 in 95 AIS cases and 282 healthy non-AIS controls from Northwest India. The genotyping was carried out on a Realtime PCR using TaqMan allele discrimination assay and the variant was found to be following Hardy Weinberg equilibrium. The statistical analyses of the genotyping data did not show significant association (p=0.66) of variant rs11190870 with AIS in the population of Northwest India. The results are interesting findings in a population that has never been studied before for AIS susceptibility. However, the findings can be attributed to under power study thus, need evaluation in a large sample set from the population. Interestingly, frequency distribution of the variant in Indian control population datasets was found to be different than other global populations. Linkage Disequilibrium (LD) differences in the genomic region were also observed in these populations while analysing 1000Genomes phase 3 data. It hints at existence of either haplotypic differences in LBX1 locus in South-Asian Indian populations with respect to other populations or genetic heterogeneity in AIS susceptibility. This lays a foundation for genome wide association study (GWAS) in Indian populations cohort, for better understanding of AIS, a task we are pursuing.